Papers with online environments
Outcome-Constrained Large Language Models for Countering Hate Speech (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven. |
| Approach: | They develop automatic counterspeech generation methods that incorporate two desired conversation outcomes into the text generation process: low conversation incivility and non-hateful hater reentry. |
| Outcome: | The proposed methods incorporate two desired conversation outcomes: low conversation incivility and non-hateful hater reentry. |
Towards Detecting Contextual Real-Time Toxicity for In-Game Chat (2023.findings-emnlp)
Copied to clipboard
| Challenge: | ToxBuster is a simple and scalable model that reliably detects toxic content in real-time for a line of chat by including chat history and metadata. |
| Approach: | They propose a model that detects toxic content in real-time for a line of chat by including chat history and metadata. |
| Outcome: | The proposed model outperforms conventional toxicity models across popular multiplayer games including Rainbow Six Siege, For Honor, and DOTA 2 and 6% of unreported toxic players can be proactively moderated. |
Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for detoxification of text often rely on manually annotated data . xiangli: "detoxification of texts is a powerful way to remove toxic content" |
| Approach: | They propose a reinforcement learning framework that optimizes detoxification and semantic preservation without annotating large amounts of data. |
| Outcome: | The proposed method overcomes major limitations and surpasses humanannotated references across multiple benchmarks. |